Medical Image Analysis
○ Elsevier BV
Preprints posted in the last 7 days, ranked by how well they match Medical Image Analysis's content profile, based on 35 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Zheng, J.; Kalaie, S.; Ma, Q.; Meng, Q.; Rjoob, K.; Gifani, P.; Hu, L.; Babazade, N.; Coriano, M.; Zhong, W.; Vafaeezadeh, M.; Tahasildar, S.; Vadgama, N.; Senevirathne, D. S.; Santhirasekaram, A.; McGurk, K. A.; Curran, L.; He, Y.; Chen, L.; Mo, Y.; Huang, L.; Qiao, M.; Huang, Y.; Bai, W.; O'Regan, D. P.
Show abstract
Cardiac imaging enables quantitative assessment of cardiac structure and function but remains constrained by cost, infrastructure and specialist expertise. In contrast, electrocardiogram (ECG) is widely accessible yet underexploited, despite encoding latent information about cardiac physiology. Here we introduce visionECG, a conditional flow matching framework that learns a probabilistic mapping between two biological distributions - the space of cardiac electrical signals and the space of cardiac geometries. Using 71,132 paired ECG and cardiac mesh sequence datasets from the UK Biobank, with external assessment in 5,000 patients with ECG-echocardiogram pairs, the model reconstructs quantitatively accurate spatiotemporal representations of the left ventricle using ECG inputs and basic demographic information alone. These reconstructions enable discrimination of structural abnormalities and disease labels, provide visualisations of functional abnormalities, and support flexible quantification of both global and regional parameters. By reframing the ECG as a generative source of patient-specific left ventricular geometry and motion, this work establishes a scalable framework for translating low-dimensional signals into high-dimensional, physiologically grounded structured representations.
Levitis, E.; Tregidgo, H. F. J.; Zimmerman, D.; Jung, B.; Karandikar, S.; Gardner, M.; Mattisson, P.; Kafadar, E.; Zapaishchykova, A.; Kann, B. H.; Sotardi, S. T.; Vossough, A.; Huang, H.; Billot, B.; Iglesias Gonzales, J. E.; Alexander, D. C.; Alexander-Bloch, A. F.; Seidlitz, J.
Show abstract
Clinical brain MRIs from pediatric health systems represent a viable resource for modeling early neurodevelopmental trajectories and studying neurodevelopmental risk in real-world populations. However, a limitation to date has been the performance of existing segmentation tools for measuring various brain phenotypes in clinical scans. In particular, many tools underperform in infant scans due to morphological and physical changes such as rapid myelination. Here, we introduce ClinSeg: a robust segmentation approach tailored to early-life clinical MRIs with variable orientation, resolution, and contrast. We leverage existing registration and synthetic data generation tools to construct a training corpus for a 3d U-Net spanning anatomical and contrast diversity, including scans with morphological abnormalities from a pediatric hospital. Validated against manual segmentations, ClinSeg outperforms existing models in infancy while matching them in childhood and adolescence. Finally, ClinSeg enables the construction of reference brain growth trajectories in 11,699 individuals from 0-21 years of age, leading to the detection of more nuanced age-related findings in clinical groups.
dela Sotta, T.; Saavedra, J. M.; Chang, V.; Xavier, A.; Henriquez, H.; Orellana, Y.; Curimil, J.
Show abstract
Diffusion models achieve high reconstruction quality in low-dose computed tomography (LDCT), but their iterative sampling trajectories impose substantial computational costs. Unlike unconditional generation, paired LDCT reconstruction starts from an image that already contains the anatomy and spatial structure of the standard-dose CT (SDCT) target; reconstruction primarily requires correcting dose-related noise and artifacts. We therefore introduce Residual Endpoint Flow Matching (REFM), an LDCT reconstruction method that learns to transport an LDCT image directly toward its paired SDCT endpoint rather than defining a noise-to-image trajectory. REFM predicts the residual velocity along linear interpolations between both images and supports single-step and multi-step reconstruction using the same trained network. We evaluate five model capacities using 1 to 50 Euler steps against deterministic U-Net and diffusion-based baselines. Across all REFM variants, one-step inference consistently provides the highest reconstruction quality. On the TCIA validation set, REFM Base achieves 50.98 dB PSNR and 0.9865 SSIM at 94.54 fps, compared with 50.92 dB, 0.9847, and 9.26 fps for DDPM-10. REFM Small retains 50.71 dB while increasing throughput to 198.56 fps. Without fine-tuning, REFM Base also matches the 25-step DDPM baseline on the external Mayo Clinic dataset, although DDPM remains stronger on synthetically degraded CRLM images. Thus, our results show that exploiting paired anatomical correspondence enables diffusion-level LDCT reconstruction with a single step reconstruction.
Lu, Z.; Uddin, S.; Uribe, S.; White, S.; Martins, R. T.; Chau, S.; Mosaddek, A. S. M.; Islam, M. S.; Nahar, N.; Azad, A. K. M.; Hossain, K. M. N.; Choudhury, H. S.; Hasan, K. M. R.; Mosaddek, N.; Rahman, S.; Hossain, M. M.; Sizar, K. M. M. H.; Angione, C.; Lio, P.; Islam, M. T.; Moni, M. A.
Show abstract
Stroke remains a leading cause of mortality and long-term disability worldwide, yet rapid diagnosis is often limited by the shortage of trained radiologists, particularly in resource-constrained settings. Automated analysis of CT imaging offers a potential solution, but existing methods often struggle to achieve clinically generalisable performance while jointly addressing multiple diagnostic tasks. Here we present the Intelligent Integrated Stroke Diagnosis System IISDS, an end-to-end deep learning framework built upon StrokeGNN, a graph-based architecture that integrates 3D contextual feature extraction with U-Net-based 2D lesion segmentation to enable comprehensive stroke analysis from non-contrast CT scans. IISDS performs stroke subtype classification, lesion segmentation and lesion volume estimation within a unified pipeline. To develop and validate the system, we collected and curated BGD-ISD through a collaboration between AI researchers, neurologists, radiologists and clinicians, resulting in a large multi-centre dataset comprising 1,507 CT scans from 597 stroke cases acquired across six hospitals and medical centres in Bangladesh. Across BGD-ISD and multiple publicly available datasets, IISDS achieves state-of-the-art performance on all tasks, improving segmentation accuracy by [≥]0.011 Dice score, reducing lesion volume estimation error by [≥]0.3 average symmetric surface distance (ASSD), and increasing classification performance by [≥]0.018 area under the receiver operating characteristic curve (AUC) compared with existing approaches. These results demonstrate the potential of graph-based deep learning to enable clinically generalisable, automated and scalable stroke diagnosis from CT imaging, supporting rapid clinical decision-making, particularly in healthcare environments with limited access to expert radiological interpretation.
Leibovici, A.; Espinos Soler, E.; Mesika, D.; Tsarfaty, G.; Livny, A.; De Santis, S.; Eggl, M. F.
Show abstract
Diffusion-weighted MRI, beyond the commonly used diffusion tensor framework, offers a unique window into tissue microstructure in vivo, yet its clinical adoption has remained limited. Major barriers include the complexity of diffusion MRI sequence design, lengthy acquisition protocols, and the challenges associated with robust estimation of high-dimensional microstructural model parameters. Here, we address these limitations by combining optimised diffusion encoding with state-of-the-art simulation-based inference, establishing a clinically feasible framework for multi-compartment diffusion modelling. We validate the approach through i) in-depth in silico experiments and ii) in vivo studies made up of both human and rodent data. The resulting microstructural metrics are robust, reproducible across healthy individuals and show significant spatial associations with brain-wide expression patterns of cell-specific genes. Requiring less than 10 minutes of acquisition time, this framework substantially lowers the barriers to advanced microstructural imaging, a prerequisite step toward its eventual evaluation for the diagnosis, stratification, and monitoring of brain disorders.
Chau, G. N.; Biswas, B. A.; Wagle, B. R.; Maeder, M. E.; Yu, J. B.; Bhattacharya, I.
Show abstract
Automated lesion segmentation is increasingly central to PSMA PET/CT interpretation, supporting staging, treatment planning, and response assessment at a scale that outpaces available nuclear-medicine expertise. However, automated PSMA-PET/CT whole-body lesion segmentation models are trained on images alone, with no knowledge of where in the body prostate metastases actually tend to occur. Radiologists use clinical domain knowledge of metastatic spread, but its absence in machine learning models produces false positives in anatomically implausible locations and missed lesions in high-risk sites such as the liver. In this work, we explore whether population-level spatial knowledge of metastatic spread can be used to augment deep learning segmentation predictions, and how such a prior should be fused with a network's output, without additional training. We build a data-driven metastasis atlas from 375 expert-annotated whole-body PSMA PET/CT scans and investigate its fusion with a trained segmentation network under a Bayesian framework, in which prediction probabilities from an nnU-Net-based lesion segmentation model serve as the likelihood and the data-driven atlas as the prior. Because metastases occupy only a small fraction of whole-body voxels, the atlas's peak probability is too low, and standard power-scaled or naive Bayesian pooling references lack the tools to deal with this shortcoming. This causes these standard fusion strategies to fail and, in the naive Bayesian case, to sharply degrade performance. We instead derive a calibrated, background-referenced log-odds fusion, one of many possible approaches to combine a population atlas with a deep learning model's predictions, distinct from classical multi-atlas label fusion in that it fuses a single population prior with a trained network's softmax rather than combining several registered atlases. Furthermore, this approach is neutral outside atlas support by construction, reduces exactly to the baseline network when unweighted, and requires no retraining. This atlas fusion significantly improved mean Dice over the baseline nnU-Net on a disjoint internal test set ($+0.011$, Holm-adjusted $p=0.021$) and on an independent external cohort ($+0.0129$, Holm-adjusted $p=3.8\times10^{-16}$), with lesion sensitivity improving from 0.849 to 0.861 internally and Dice improving over baseline in every stratified anatomic region, including the rare, high-risk sites motivating this work, while naive Bayesian pooling degrades performance sharply and power-scaled pooling underperforms it throughout. Our findings suggest that population-level spatial priors can meaningfully augment deep learning predictions in whole-body oncologic segmentation, provided the fusion rule is calibrated to where the prior actually carries signal.
Aicher, A.; Graf, R.; Kirschke, J.; Frauenfelder, T.; Ensle, F.; Menze, B.; Decker, J.; Kröncke, T.; Haubold, J.; Ringhof, S.; Bamberg, F.; Schmidt, C. O.; Wielpütz, M.; Leitzmann, M.; Willich, S. N.; Keil, T.; Niendorf, T.; Pischon, T.; Schlett, C.; Möller, H.
Show abstract
Rib-cage morphology is a determinant of thoracic biomechanics, ventilation, and injury response, yet statistical shape models (SSMs) of the rib cage have relied on small cohorts (~100s of individuals) imaged by clinical computed tomography, which over-represents injury and disease. We constructed a surface-based SSM of the complete 24-rib cage from 26,275 standardised whole-body magnetic resonance imaging (MRI) scans of adults aged 19-74 years from the population-based German National Cohort (NAKO). Ribs were segmented with a deep-learning pipeline (a rib-extended SPINEPS model), reconstructed as per-rib surface meshes, and brought into dense vertex-wise correspondence by Gaussian-process morphable registration in Scalismo; the aligned ensemble was summarised by generalised Procrustes analysis and principal component analysis (PCA). Fourteen per-rib geometric descriptors provided a quantitative cross-walk between the abstract PCA modes and named shape features, and associations with sex, age, body size and composition (including body-fat percentage), and smoking exposure were estimated by multivariable regression with Benjamini-Hochberg false-discovery-rate control. Shape variation was strongly concentrated: 28 modes captured 95% of the total variance, and the first three alone accounted for 69.4% (PC1, 42.6%; PC2, 16.3%; PC3, 10.5%) and admitted consistent anatomical readings - a sexually dimorphic axis (PC1), a slender-versus-stout body-habitus contrast (PC2), and a free-rib-size axis at ribs 11-12 (PC3). The sexes were nearly fully separated along PC1 (Cohen's d = 2.52). Body mass and body-fat percentage were the dominant modifiable correlates of rib-cage shape, whereas the association with cumulative smoking exposure was comparatively small. The model is released as a population-representative geometric reference for benchmarking and morphing donor-derived finite-element human-body models and for further large-cohort shape analysis.
Ye, Z.; He, F.; Zhao, T.; Xia, W.
Show abstract
Ultrathin endoscopy is highly attractive for real-time tissue imaging in narrow and hard-to-reach regions of the body. A single multimode fibre (MMF) is an attractive probe because of its small diameter, flexibility, and diffraction-limited spatial resolution enabled by the large number of transverse modes guided within a single core. Because the distal fibre tip is inaccessible during endoscopy, reflection-mode imaging, in which the same fibre delivers illumination and collects backscattered light, is more practical than transmission-mode imaging. However, image recovery from the resulting speckle pattern is challenging because light undergoes double-pass propagation through the MMF, with mode coupling and dispersion; the backscattered signal is weak, and the camera records intensity only, without phase information. Here, we propose a single-shot reflection-mode MMF imaging framework that combines a reflected real-valued intensity transmission matrix (reflected-RVITM) with an image restoration network. The reflected-RVITM is calibrated using intensity-only measurements, without interferometry or phase retrieval, and provides a physics-guided initial reconstruction from a single backscattered speckle frame. A restoration network then refines this initial reconstruction instead of inverting the raw speckle. Four restoration backbones are evaluated: HPM-Attention-UNet, GAM, MambaIRv2, and CICPNet. On matched datasets, hybrid models outperformed corresponding networks trained to map raw speckle directly to images. For example, HPM-Attention-UNet on MNIST improved mean PCC from 0.572 to 0.944 (+65.1%). Under domain shift, with training only on Fashion-MNIST and tested on unseen CIFAR scenes, hybrid models achieved mean PCC of 0.61-0.65, compared with 0.36-0.50 for direct learning. This framework is further demonstrated using physical objects at the distal fibre tip. These results demonstrate that a reflected-RVITM physics prior combined with a restoration network enables single-shot image recovery after intensity-only calibration, offering a phase-retrieval-free and generalisable route towards minimally invasive reflection-mode MMF endoscopy.
Zhuang, H.; Zakama, A.; Heller, K.; Faulkner, S.; Gollub, B.; Young-Lin, N.; Chen, I. Y.; Asiedu, M.
Show abstract
In this work, we demonstrate the unprecedented value of NIH's "All of Us Research Program" (AoURP) dataset in studying maternal morbidity and building predictive machine learning (ML) models across heterogeneous populations in the United States. We developed robust and data-driven preprocessing pipelines to curate a longitudinal, multi-site, multimodal, and demographically diverse pregnancy dataset (20,253 subjects; 27,525 pregnancy episodes) from AoURP data, using electronic health records (EHR) (Conditions, Labs, Measurements) and survey responses (Social Determinant of Health (SDoH)), focusing on 7 crucial maternal health adverse outcomes. After characterizing data quality, missingness, and heterogeneity, we performed statistical correlation analysis to identify risk factors. We subsequently developed XGBoost and sequential LSTM models to predict the adverse outcomes, reaching state-of-the-art performance for multiple outcomes. We conducted model interpretability post-hoc analysis to understand success points and fairness analysis to evaluate implications for socio-economic disparities. Four practicing physicians reviewed the set of statistically significant and ML model identified features to assess their clinical validity and novelty. Most features identified through either statistical correlations or ML feature importance analysis aligned with known clinical risk factors. Several features were identified that the ML models used but that are not currently used in clinical practice and may merit further clinical investigation. Fairness analysis revealed certain associations with SDoH and age highlight areas that warrant continued monitoring. Overall, we demonstrate that meaningful populational level patterns can be extracted, and high-performing machine learning models can be trained on this longitudinal, diverse, multi-site dataset. Important risk features, particularly novel ones identified, if validated, could inform new strategies for maternal care or enable development and validation of outcome-specific, clinically deployable ML models.
Zeng, H.; Hu, M.; Phng, L.-K.; Matsunaga, Y. T.
Show abstract
Three-dimensional (3D) mural cell morphology is heterogeneous and coupled to vessel geometry, however, measurements from two-dimensional (2D) maximum intensity projections (MIP) obscure overlapping processes and cell-vessel contacts. Accordingly, we developed Mural-VISTA, a semi-automated Python workflow for mural cell-vessel interaction and single-cell topo-morphology analysis of reconstructed surface meshes. This workflow integrates mesh pretreatment, interactive centerline extraction, hierarchical segmentation of cell soma, main axis and secondary processes (branches), and extraction of 36 multiscale (cell process segment level, process level, and whole cell level) topo-morphological and vessel-referenced metrics. Mural-VISTA identified morphological changes in pericytes and vascular smooth muscle cells (vSMCs) with altered RhoA activity. Constitutive active RhoA (RhoA CA) over-expression reduced branch complexity and increased process alignment in both cell types, while increased whole-cell and branch solidity only in vSMCs. Dominant negative RhoA (RhoA DN) over-expression increased branch abundance and reduced branch solidity in pericytes but not vSMCs, suggesting cell-type specific effect of reduced RhoA activity. In conclusion, Mural-VISTA enables quantitative 3D profiling of mural cell architecture and its spatial relationship with the vessel.
Buzzanca, G.; Pala, C.; He, J.; Hofstraat-Boersma, R.; Tammaro, A.; van Midden, D.; Buelow, R.; Hoelscher, D. L.; Muehlfeld, A. S.; Koeller, m.; Kozakowski, N.; Boehmig, G.; Halloran, P. F.; van der Helm, D.; Meziyerh, S.; Venhuizen, J.-H.; Haitjema, S.; Dijkstra, J.; Hilbrands, L. B.; Steenbergen, E. J.; van Zuilen, A. D.; Nurmohamed, A. S.; Bemelman, F. J.; Bruns, I. B.; Callegaro, G.; van de Water, B.; Pieters, T. T.; Breimer, G. E.; Rossi, G. M.; Fiaccadori, E.; Maggiore, U.; Roelofs, J. J. T. H.; Testa, F.; Fontana, F.; Abiola, A. A.; Delsante, M.; Corthals, G. L.; Peters-Sengers, H.; Ngu
Show abstract
Accurate, reproducible interpretation of kidney allograft biopsies is critical for diagnosis of graft injury to guide prognosis and management. The international Banff classification is a consensus diagnostic system based on semiquantitative histological lesion scoring on either extent or severity of kidney transplant biopsies. However, pathologist scoring is limited by substantial interobserver variability, constrained scalability, and the inherent nature of the scoring system itself. Here we present BanffNET, a weakly supervised, probabilistic deep learning framework that combines self-supervised feature extraction with a novel Bayesian multiple-instance learning framework to predict (continuously) the full spectrum of Banff lesion scores directly from whole-slide images (WSIs). Using lesion-specific aggregation functions tailored to localized (modeling lesion severity) and diffuse pathologies (modeling lesion extent), BanffNET generates interpretable, patch-level probability maps and calibrated slide-level scores. BanffNET's performance was assessed relative to consensus, biological correlates of rejection and clinical outcome, demonstrating superior consistency, transportability and generalization. Trained on 7,249 WSIs from three cohorts, BanffNET demonstrates consistent performance on 11,028 WSIs across five external test sets, performing on par or exceeding expert consensus across lesions. BanffNET scores align more closely than pathologist Banff scores with molecular profiles of rejection, offering a transparent, biologically grounded framework for computational pathology with relevance beyond transplantation.
Tecchio, P.; Schlaffke, L.; Bolsterlee, B.; Hahn, D.; Raiteri, B. J.
Show abstract
Muscle architecture shapes muscle function and changes with age, growth, training and disease, yet quantifying three-dimensional (3D) muscle architecture in vivo remains challenging. We introduce a hybrid fascicle tractography approach for freehand 3D ultrasound data that accurately reconstructs 3D muscle fascicles with respect to an objective, anatomically relevant coordinate system defined by the muscle's central aponeurosis. The hybrid approach combines Hessian-based fascicle detection with wavelet-based refinement to generate volumetric fascicle orientations. In a synthetic dataset with known ground truth, fascicle orientations and lengths were estimated with errors of [≤]2{degrees} and ~1.5%, respectively. In vivo, the approach detected physiologically plausible fascicle lengthening in the human tibialis anterior following a passive plantar flexion rotation, whereas diffusion tensor imaging of the same muscle did not. The proposed method enables anatomically relevant, objective and non-invasive quantification of 3D muscle architecture in vivo, providing a practical framework for applications in clinical and applied muscle physiology.
Shi, Z.; Budhkar, A.; Amin, W.; Pollok, K. E.; Su, J.; Huang, K.
Show abstract
Improvements in data availability, sharing, and integration, together with the development of explainable artificial intelligence (XAI) techniques, are advancing precision medicine for pediatric cancer by facilitating diagnosis, biomarker discovery, and drug development. Data sharing commons and initiatives like the Childhood Cancer Data Initiative (CCDI) provide access to pediatric-specific genomic and clinical data cohorts and improve data availability for pediatric cancer research. Based on CCDI, a scalable AI platform, Graph Artificial Intelligence for Pediatric Oncology (GAIPO), integrates various data modalities from bulk and single-cell omics data to clinical information. Such multi-modal data facilitates the training and development of advanced XAI models for pediatric cancers. We then developed an end-to-end multi-modality framework, PCGS, for pediatric cancer by incorporating omics-specific representation learning via GNN models with cross-attention fusion and multi-objective learning for downstream tasks such as classification, clustering, and survival analysis. This framework outperforms previous supervised multi-omics integration baseline approaches based on glioma and Wilms tumor cohorts and enables GNN model explainability via Shapley value-based feature attribution approaches to explain the contributions of gene-level features across various biomedical tasks, including classification and survival. Given specific background samples (e.g., age groups, sex, grades) as baselines, this explainable GNN model estimates and ranks the importance scores for input features from each omics modality. It identifies background-specific key features for biomarker discovery, risk group identification, and survival analysis in glioma and Wilms tumor, with potential applicability to other pediatric cancers.
Perlman, A.; Goldstein, N.; Goldman, M.; Shapiro, M.; Barash, E.; Bar, A.; Raveh, T.; Tordjman, E.; Schussheim, H.; Dormont, F.; Matalon, O.
Show abstract
Background. Cardiovascular-outcomes trials are lengthy, costly, and associated with substantial uncertainty prior to readout. In-silico trial simulation using real-world data (RWD) has emerged as a potential tool to support earlier decision-making; however, evidence of prospective predictive validity, generated prior to trial result disclosure, remains limited. Methods. We applied a semi-mechanistic machine learning framework integrating real-world patient data with biologically informed drug representations to prospectively simulate the VESALIUS-CV trial evaluating evolocumab versus placebo. The simulation model was trained on a combination of patient-level real-world data and a drug-centric knowledge graph and validated for both patient-level and trial-level retrospective predictive performance. The model was then used to simulate VESALIUS-CV before public disclosure of trial results, using a locked model and prespecified eligibility criteria and primary endpoint aligned with the clinical protocol. A patient-level time-to-event model was used to generate virtual trial arms, from which cumulative incidence curves, hazard ratios, confidence intervals, and p-values for major adverse cardiovascular events (MACE) were estimated. Results. In retrospective validation, the model demonstrated strong patient-level discrimination, with time-dependent ROC-AUC values ranging from 0.80 to 0.90 across follow-up horizons. For trial-level validation, 22 randomized cardiovascular-outcomes trials were simulated, and hazard ratios for 3-point MACE across 24 between-arm comparisons showed consistent directional agreement and quantitative correlation with published results such that the model accurately predicted trial success, achieving an F1 score of 0.83, with precision of 0.79 and sensitivity of 0.89. In a fully prospective application, the simulation predicted a statistically significant reduction in 3-point MACE with evolocumab versus placebo, estimating a hazard ratio of 0.78 (95% CI, 0.70-0.87) at 54 months. These predictions were consistent with the subsequently reported VESALIUS-CV results, which demonstrated a hazard ratio of 0.75 (95% CI, 0.65-0.86) at 55 months of median follow-up. Conclusions. In a fully prospective setting, a RWD-driven, AI-based simulation accurately predicted the direction, magnitude, and temporal dynamics of treatment effects observed in the VESALIUS-CV trial. These results demonstrate that in-silico trial simulation can anticipate clinical outcomes in the prospective setting, supporting its use as a complementary tool for early decision-making, trial design optimization, and de-risking in cardiovascular drug development.
Oyarzun Silva, R.; Hernandez Hernandez, P.
Show abstract
Background. Accurate delineation of the gross tumour volume (GTV) - primary tumour (GTVp) and nodal disease (GTVn) - on FDG-PET/CT is a critical step of head and neck radiotherapy planning. Comparisons between lightweight custom networks and the auto-configured nnU-Net v2 are usually reported as end-to-end pipelines, conflating the contribution of the network with that of the inference-time post-processing applied on top of it. We separated the two. Methods. MiniUNet3D (custom 3D U-Net, 18.3 M parameters) and nnU-Net v2 (3d_fullres, 88.2 M parameters) were trained on the same 578 FDG-PET/CT cases (85/15 author-defined split of the HECKTOR 2025 Task 1 set, 8 centres) and evaluated on the same internal cohort. Three arms were compared pairwise: MiniUNet3D raw output at a fixed 0.5 threshold, MiniUNet3D with a locked adaptive post-processing pipeline, and nnU-Net v2. Comparisons used paired Wilcoxon tests with bootstrap confidence intervals, Bonferroni and Benjamini-Hochberg correction, and Cohen's d; catastrophic failure (Dice < 0.01) was compared with an exact McNemar test. Cases with an empty reference for a given target were excluded from that target's analysis (n = 98 GTVp, n = 93 GTVn). Results. With post-processing matched off, nnU-Net v2 was superior: median GTVp Dice 0.799 versus 0.592 (mean difference -0.244, 95 % CI -0.300 to -0.191; d = -0.88) and GTVn 0.774 versus 0.598 (d = -0.82). Post-processing raised MiniUNet3D to 0.800 (GTVp) and 0.738 (GTVn), recovering 79 % of that difference. Post-processed, MiniUNet3D matched nnU-Net v2 on GTVp Dice (p = 0.113) but remained inferior on nodal disease after Bonferroni correction (Dice p = 0.041; surface Dice p = 0.049). Catastrophic GTVp failures were 25/98 raw, 8/98 post-processed and 1/98 for nnU-Net v2 (McNemar p = 0.016). Inference took 34 s versus 78 s per case on the same GPU. Conclusions. Post-processing recovered most, but not all, of the difference between the two models, and it did not confer robustness: an eight-fold higher rate of empty contours on small primaries persisted, which is the more consequential difference for planning safety. Pipeline comparisons reported without a post-processing ablation risk attributing to a network what post-processing supplied.
Pandey, D.; Narasimhan, V. M.
Show abstract
Self-supervised models increasingly convert medical images into quantitative phenotypes for biological discovery, but statistical reproducibility does not establish that a learned phenotype represents the intended anatomy. We trained a video masked-autoencoder on 69,932 UK Biobank cardiac cine-MRI studies and performed genome-wide association analysis of its latent representation. Although 18 of 20 leading axes were heritable with well-calibrated statistics, the representation encoded substantial field-of-view information: body size, stature and imaging centre (linear-probe R^2=0.55 for site); standard genomic-control and LD-score diagnostics did not identify this source of phenotype-level confounding. Restricting the field of view to the heart and residualising body and acquisition covariates before dimensionality reduction substantially attenuated linear and non-linear nuisance information while retaining cardiac signal. Adjusting the same covariates only during association testing attenuated nuisance associations but recovered substantially less of the cardiac-associated genetic signal, consistent with nuisance variation having already influenced the principal-component basis. The corrected representation identified new associated loci beyond those detected using supervised phenotypes at matched sample size, which shared genetic architecture selectively with cardiac-conduction traits and were localised to cardiac structures within the imaged field of view. Confounding in learned medical phenotypes can arise upstream of association testing, highlighting the importance of auditing and, where appropriate, correcting learned representations before association testing.
Courtens, J.; Muller, F. M.; Li, E. J.; Vanhove, C.; Vandenberghe, S.; Pantel, A. R.; Karp, J. S.; Daube-Witherspoon, M. E.
Show abstract
Dynamic positron emission tomography (PET) with long axial field-of-view (LAFOV) scanners enables multi-organ imaging and kinetic quantification beyond static (late-phase) imaging; however, the long times typically required for dynamic acquisitions remain clinically impractical. This study evaluates a deep learning (DL) framework to enable abbreviated dynamic PET acquisitions, comparing single-time-window (STW, early dynamic data only) and dual-time-window (DTW, early dynamic data plus a late 5-min static frame) protocols with early dynamic scan durations of 5-30 min and dose levels ranging from 360 MBq to 18 MBq. Seventeen 60-min dynamic [18F]FDG datasets were first motion-corrected using a staggered FALCON pipeline and then used to train and test a spatiotemporal DL model for autoregressive frame prediction. Performance was assessed across the full quantitative workflow, from DL-predicted frames and time-activity curves to organ-based kinetic modeling and voxel-wise parametric imaging in multiple tissues and two patient cohorts. DTW protocols consistently outperformed STW, better preserving late-phase kinetics. For a 15-min early dynamic scan, adding a late 5-min scan reduced mean absolute Ki difference from 23% (STW) to 17% (DTW) in the liver and from 26% to 15% in the thalamus. DTW + DL further reduced errors to [≤]10% in the liver, thalamus, and breast lesion, and 16% in muscle. Our recommended protocol, 15-min early dynamic scan plus a 5-min late scan with DL, remained robust to up to a 5-fold dose reduction (~74 MBq). Overall, these findings support DL-enabled abbreviated, low-dose dynamic LAFOV PET as a clinically feasible approach for accurate kinetic quantification
Kim, Y.; Heo, W.; Park, S. J.; Kim, Y.; Cho, Y. E.
Show abstract
Molecular staging of Alzheimer's disease (AD) increasingly defines transition boundaries along single-cell pseudo-progression trajectories, yet whether such boundaries reproduce across brain regions, cohorts and molecular modalities is rarely tested. We present a permutation-controlled audit that combines nine boundary-detection algorithms with a fixed marker panel and four orthogonal reproducibility axes-algorithmic consensus, region, cohort and modality. On synthetic data with planted ground-truth boundaries the audit reaches 100% sensitivity and 94% specificity, rejecting four distinct artefact classes each by a different axis. Applied to the Seattle Alzheimer's Disease Brain Cell Atlas middle temporal gyrus, it localizes a transition that is robust across algorithms and recovered in most cell types but does not generalize: its leading marker is attenuated or absent in prefrontal cortex, entorhinal cortex and cerebrospinal fluid, and an apparent cross-region conservation of glial metabolic genes proves to be a global-expression offset rather than a shared program. The same audit nonetheless certifies an externally validated marker (astrocytic PTGDS) as reproducible across regions and modalities, showing that it separates generalizable anchors from dataset-specific ones rather than rejecting all signals. We provide this four-axis audit as a transferable, code-available standard to apply before a trajectory boundary is read as a biological stage, in AD and other progressive proteinopathies.
Ekambarapu, L.; Pendyal, A.; Lin, A.; Alwakeel, M.; Rajaratnam, A.
Show abstract
Background: Unstructured biomedical data, such as echocardiography reports, are rich in information but time consuming to analyze at scale. Rule-based, regular expression-driven terminology mapping can only extract individual variables while large language models (LLMs) offer scalable and clinically meaningful interpretations of heterogeneous disease processes. Right ventricular dysfunction (RVD) is an example of a multifactorial disease state in which key structural and physiologic features are captured both narratively and in structured fields, making it an ideal test case for evaluating whether LLMs can recover complex phenotypes that rules based methods routinely miss. Purpose: To compare an LLM-based extraction method to a conventional rules-based schema for identifying and phenotyping echocardiographic features associated with RVD in a large TTE dataset. Methods: MIMIC-III NOTE2NUM echocardiography reports (n = 45,794) were analyzed using GPT-4o-based LLM extraction deployed within a secure health system enclave and were benchmarked against echocardiographic measurements defined in the MIMIC-III dictionary schema. In MIMIC-III, PH was recorded qualitatively (mild/moderate/severe) based on tricuspid regurgitant (TR) jet velocity and then re-coded as present vs. absent. LLM based extraction defined RVD as (1) RV structural abnormality (>= 1 of hypertrophy, dilation, or wall hypo-/akinesis) or (2) RV pressure/volume overload (>= 2 of the following: estimated right atrial pressure > 8 mmHg, TR jet velocity > 2.8 m/s, fractional area change < 35%, tricuspid annular planar systolic excursion < 17 mm, S' < 9.5 cm/s, or E/e' > 14), with PH defined as estimated pulmonary artery systolic pressure > 35 mmHg or qualitative documentation of PH. Results: LLM extraction identified PH in 15,394 (33.6%), RV pressure/volume overload in 14,449 (31.6%), and RV structural abnormalities in 11,955 (26.1%). Co-occurrence was common: overload + structural changes in 9,380 (20.5%), overload + PH in 9,756 (21.3%), structural changes + PH in 6,183 (13.5%), and all three in 5,620 (12.3%). Using the MIMIC-III dictionary schema, PH prevalence was similar (15,371; 33.6%), but RV overload fields were captured less often (pressure overload 1,357 [3.0%], volume overload 1,128 [2.5%], pressure + volume overload 1,093 [2.4%]; any overload field 3,578 [7.8%]), and RV pressure/volume overload with PH was identified in only 731 (1.6%). Conclusions: LLM-based extraction outperforms rules-based schemas for identifying complex disease states not defined by any single variable. By synthesizing multifactorial signals, LLMs can phenotype RVD with higher fidelity and support population-level assessment. Further validation using multimodality imaging, invasive hemodynamics, and clinical outcome data is needed.
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.